TurboFieldfare runs Gemma 4 on low RAM Macs

21 09 2026

TurboFieldfare is a Swift and Metal runtime that runs Gemma 4 26B-A4B using roughly two gigabytes of RAM. It streams expert weights from the SSD to keep the model in memory, allowing it to run on an eight gigabyte M2 MacBook Air. The project includes a native Mac app, a CLI, and an OpenAI-compatible server, all built specifically for Apple Silicon.

It is worth trying if you need local inference on low-memory hardware and want to avoid the overhead of general-purpose wrappers. The catch is that decode speed on base M2 hardware is slow, ranging from five to six tokens per second. This tool is not suitable for high-throughput tasks or older Intel Macs.

more: https://github.com/drumih/turbo-fieldfare


Actions

Information

Leave a comment




Design a site like this with WordPress.com
Get started